Skip to main content

Open Data

Open data is data anyone can access, use and redistribute, subject at most to attribution and share-alike. In health, the useful question is not whether to be open but which data can be opened, at what granularity, and with what safeguards.


What "open" requires​

  1. Available — retrievable in bulk, at no more than reproduction cost
  2. Machine-readable — CSV, JSON or Parquet, not a PDF of a table
  3. Openly licensed — explicit reuse rights, stated on the dataset
  4. Findable — catalogued, with a stable URL
  5. Documented — a data dictionary, not just column headers

A spreadsheet on a ministry website with no licence and no schema meets the letter of "published" and none of the intent.


What can be opened​

Usually safe to open:

  • Facility registries — locations, services, ownership
  • Aggregate service statistics above a defined suppression threshold
  • Health workforce counts by cadre and administrative level
  • Commodity availability and stock-out rates
  • Budget and expenditure data
  • Metadata: indicator definitions, code lists, reporting calendars

Requires careful handling or should not be opened:

  • Any individual-level clinical record
  • Small-area data on stigmatised conditions
  • Data that identifies individual providers by outcome
  • Cells small enough to identify a person

Disclosure risk​

Aggregate does not mean anonymous. Risks to manage:

  • Small cells. Adopt a suppression threshold and apply it consistently.
  • Differencing. Two publications that overlap can reveal a suppressed cell by subtraction; suppress secondary cells too.
  • Linkage. Combining an open dataset with another source can re-identify individuals even when neither does alone.

See health data for de-identification technique and data governance for who authorises release.


Publishing well​

  • Stable identifiers. Facility codes that change between releases break every downstream user.
  • Versioned releases with a changelog; never silently overwrite.
  • A data dictionary giving definition, unit, period and source for each field.
  • Documented provenance — which system, extracted when, covering what.
  • Known limitations stated up front. Users will find them anyway; stating them preserves trust.
  • An API as well as bulk files, when there is demand for both.